Papers with stack-basedattention mechanism
A Transformer with Stack Attention (2024.findings-naacl)
Copied to clipboard
| Challenge: | Recent research suggests that transformer-based language models fail to learn basic algorithmic patterns. |
| Approach: | They propose to augment transformer-based language models with a differentiable stack-based attention mechanism that adds a level of interpretability to the model. |
| Outcome: | The proposed model can model some, but not all, deterministic context-freelanguages. |